Back

Frontiers in Genetics

Frontiers Media SA

All preprints, ranked by how well they match Frontiers in Genetics's content profile, based on 230 papers previously published here. The average preprint has a 0.18% match score for this journal, so anything above that is already an above-average fit. Older preprints may already have been published elsewhere.

1
Imputation and polygenic score performances of human genotyping arrays in diverse populations

Nguyen, D. T.; Tran, T.; Tran, M.; Tran, K.; Pham, D.; Duong, N. T.; Nguyen, Q.; Vo, N. S.

2022-06-15 genomics 10.1101/2022.06.14.496059 medRxiv
Top 0.1%
45.0%
Show abstract

Regardless of the overwhelming use of next-generation sequencing technologies, microarray-based genotyping combined with the imputation of untyped variants remains a cost-effective means to interrogate genetic variations across the human genome. This technology is widely used in genome-wide association studies (GWAS) at bio-bank scales, and more recently, in polygenic score (PGS) analysis to predict and to stratify disease risk. Over the last decade, human genotyping arrays have undergone a tremendous growth in both number, and content making a comprehensive evaluation of their performances became more important. Here, we performed a comprehensive performance assessment for 23 available human genotyping arrays in 6 ancestry groups using diverse public, and in-house datasets. The analyses focus on performance estimation of derived imputation (in terms of accuracy and coverage) and PGS (in term of concordance to PGS estimated from whole genome sequencing data) in three different traits and diseases. We found that the arrays with a higher number of SNPs are not necessarily the ones with higher imputation performance, but the arrays that are well-optimized for the targeted population could provide very good imputation performance. In addition, PGS estimated by imputed SNP array data is highly correlated to PGS estimated by whole genome sequencing data in most of cases. When optimal arrays are used, the correlations of key PGS metrics between two types of data can be higher than 0.97, but interestingly, arrays with high density can result in lower PGS performance. Our results suggest the importance of properly selecting a suitable genotyping array for PGS applications. Finally, we developed a web tool that provide interactive analyses of tag SNP contents and imputation performance based on population and genomic regions of interest. This study would act as a practical guide for researchers to design their genotyping arrays-based studies. The tool is available at: https://genome.vinbigdata.org/tools/saa/

2
Predicting Clinical Phenotypes by Growth Curve Modeling of Transcriptomic Signatures during Disease Progression

Akhlaghi, M.; Ghasemi, E.; Ray, M. S.; Pyne, S.

2026-01-15 bioinformatics 10.64898/2026.01.13.699292 medRxiv
Top 0.1%
39.2%
Show abstract

High-throughput gene expression data analysis has benefited from many statistical tests of differential expression across two or more groups such as t tests, ANOVA, etc. Yet, in complex transcriptomic datasets such as longitudinal or repeated measures, few studied have addressed such key issues as group effects and temporal dependency in expression profiles with a single model that is both practically effective and theoretically grounded. In this study, we used Growth Curve Model (GCM), as a generalization of MANOVA, to identify differentially expressed longitudinal profiles of genes, and thus predicted the associated clinical phenotypes, of pediatric lupus during the progressions of the disease across two different racial groups. In particular, we detected a module of histone genes which was shown to be linked with lupus.

3
High-throughput TCRB enrichment sequencing of human cord blood exhibited a distinct fetal T cell repertoire in the third trimester of pregnancy

Dong, Y.; Wei, C.; Wang, J.; Wu, X.; Zhao, Y.; Cai, Y.; Han, Y.; Wang, Y.; Li, H.; Qiao, J.; Yuan, W.

2022-09-12 genetics 10.1101/2022.09.08.506871 medRxiv
Top 0.1%
35.4%
Show abstract

Study questionWhat are the molecular characteristics during the maturation process of the human fetal immune system in the third trimester of pregnancy? Summary answerBoth the diversity and length of complementarity determining region 3 (CDR3s) in the fetal TCRB repertoire were less than those of adult CDR3s, and the fetal CDR3 length increased with gestation weeks in late pregnancy. What is known alreadyThe adaptive immune system recognizes various pathogens based on a large repertoire of T-cell receptors (TCR repertoire), but the maturation dynamics of the fetal TCR repertoire in the third trimester are largely unknown. The CDR3is the most diversified segment in the T-cell receptor {beta} chain (TCRB) that binds and recognizes the antigen. Study design, size, and durationThis was a basic research to assess the composing characteristics of TCRBs in core blood and the dynamic pattern with fetal development in the third trimester of pregnancy. Participants/materials, setting methodsHigh-throughput TCRB-enrichment sequencing was utilized to characterize the TCRB repertoire of cord blood at 24~38 weeks of gestational age (WGA) with nonpreterm fetuses and to investigate their difference compared with that of adult peripheral blood. Main results and the role of chanceCompared to the adult control, the fetal TCRB repertoire had a 4.8-fold lower number of unique CDR3s, a comparable Shannon diversity index (p=0.7387), a lower mean top clone rate (p < 0.001) and a constrictive top 1000 unique clone rates. Although all kinds of TCRBV and TCRBJ genes present in adult CDR3s were identified in fetuses, nearly half of these fragments showed a significant difference in usage. Moreover, the fetal TCRB repertoire held a shorter CDR3 length, and the CDR3 length showed a progressive increase with fetal development. Jensen-Shannon (JS) divergences of TCRBV and TCRBJ gene usage in dizygotic twins were much lower than those in unrelated pairs. In the parental-fetal pair, JS divergence of TCRBV gene usage was not obviously different, while that of TCRBJ gene usage was only slightly lower. Limitations, reasons for cautionThe sample size is limited due to the limited accessibility to cord blood in late pregnancy with healthy nonpreterm fetuses. Wider implications of the findingsOur findings reveal the unique properties of fetal TCRB repertoires in the third trimester, fill the gap in our understanding of the maturation process of prenatal fatal immunity, and deepen our understanding of the immunologically relevant problems in neonates. Study funding/competing interest(s)This work was supported by the National Natural Science Foundation of China (82171661) and Tianjin Municipal Science and Technology Special Funds for Enterprise Development (NO. 14ZXLJSY00320). The authors declare that they have no competing interests.

4
Tissue-wide scRNA-seq analysis reveals enrichment of imprinted genes in stem and endocrine cell-types in mice

Higgs, M. J.; Hill, M. J.; John, R. M.; Isles, A. R.

2023-08-07 genetics 10.1101/2023.08.05.552036 medRxiv
Top 0.1%
33.4%
Show abstract

Enriched expression of imprinted genes may provide evidence of convergent function. Here we interrogated five single-cell RNA sequencing datasets to identify imprinted gene over-representation in the embryonic and adult mouse focusing on tissues including the bladder, pancreas, mammary gland and muscle. We identify a consistent enrichment of imprinted genes in stromal cell and mesenchymal stem cell populations across these tissues, suggesting a role in tissue maintenance. Furthermore, we identify a distinct enrichment in the endocrine islets of the mouse pancreas, over and above the stromal/stem cells from this tissue. Taken together with our previous work examining imprinted gene expression in cell subpopulations of the adult mouse brain and pituitary gland, these data suggest that genomic imprinting influences physiology largely via separate systems of cell populations either involved in hormonal signalling or in stemness and cell-fate co-ordination.

5
Multiethnic Polygenic Risk Prediction in Diverse Populations through Transfer Learning

Tian, P.; Chan, T. H.; Wang, Y.-F.; Yang, W.; Yin, G.; Zhang, Y. D.

2022-06-11 genetics 10.1101/2022.03.30.486333 medRxiv
Top 0.1%
30.4%
Show abstract

Polygenic risk scores (PRS) leverage the genetic contribution of an individuals genotype to a complex trait by estimating disease risk. Traditional PRS prediction methods are predominantly for European population. The accuracy of PRS prediction in non-European populations is diminished due to much smaller sample size of genome-wide association studies (GWAS). In this article, we introduced a novel method to construct PRS for non-European populations, abbreviated as TL-Multi, by conducting transfer learning framework to learn useful knowledge from European population to correct the bias for non-European populations. We considered non-European GWAS data as the target data and European GWAS data as the informative auxiliary data. TL-Multi borrows useful information from the auxiliary data to improve the learning accuracy of the target data while preserving the efficiency and accuracy. To demonstrate the practical applicability of the proposed method, we applied TL-Multi to predict the risk of systemic lupus erythematosus (SLE) in Asian population and the risk of asthma in Indian population by borrowing information from European population. TL-Multi achieved better prediction accuracy than the competing methods including Lassosum and meta-analysis in both simulations and real applications.

6
Multi-phase, multi-ethnic GWAS uncovers putative loci in predisposition to human sprint performance, health and disease

Wang, G.; Fuku, N.; Miyamoto-Mikami, E.; Tanaka, M.; Miyachi, M.; Murakami, H.; Cheng, Y.-C.; Mitchell, B. D.; Morrison, E.; Austin, K. G.; Ahmetov, I. I.; Sportgene Research Group, ; Generozov, E. V.; Filipenko, M. L.; Gilep, A. A.; Gineviciene, V.; Moran, C. N.; Venckunas, T.; Cieszczyk, P.; Derave, W.; Papadimitriou, I.; Garton, F. C.; North, K.; Padmanabhan, S.; Pitsiladis, Y. P.

2023-12-09 sports medicine 10.1101/2023.12.08.23299720 medRxiv
Top 0.1%
29.2%
Show abstract

The genetic underpinnings of elite sprint performance remain largely elusive. For the first time, we uncovered rs10196189 (GALNT13) in the cross-ancestry, genome-wide analysis of elite sprint and power-oriented athletes and their controls from Jamaica, the USA, and Japan, and replicated this finding in two independent cohorts of elite European athletes (meta-analysis P < 5E-08). We identified statistically significant and borderline associations for cross-ancestry and ancestry specific loci in GALNT13, BOP1, HSF1, STXBP2 GRM7, MPRIP, ZFYVE28, CERS4, and ADAMTS18, predominantly expressed in the nervous and hematopoietic systems. Further, we revealed thirty-six previously uncharacterized genes associated with host defence, leukocyte migration, and cellular responses to interferon-gamma and unveiled (reprioritized) four genes, UQCRFS1, PTPN6, RALY and ZMYM4, responsible for aging, neurological conditions, and blood disorders from the elite athletic performance cohorts. Our results provide new biological insights into elite sprint performance and offer clues to the potential molecular mechanisms interlinking and operating in elite athletic performance and human health and disease.

7
Genomic Regions and Variant Analysis Reveal Candidate Genes Associated with Age at First Service in Chinese Holstein Heifers

Pan, C.; Ma, L.; Wang, L.; Zhang, Z.; Wang, A.; Bai, F.; Cai, Y.; Ma, Y.; Wang, Y.; Liu, G.; Zhu, K.; Lv, X.; Wang, X.; Jiang, Y.

2025-07-29 genetics 10.1101/2025.07.28.667120 medRxiv
Top 0.1%
27.6%
Show abstract

Understanding the genetic basis and identifying quantitative trait loci (QTL) for age at first service (AFS) is essential for improving reproductive efficiency and reducing economic costs in heifers rearing. We conducted a genome-wide association study (GWAS) for AFS using the genomic estimated breeding values (GEBVs) of 3,686 Chinese Holstein heifers, genotyped with 7,964,217 single nucleotide polymorphisms (SNPs) after quality control. Significant SNPs and candidate genes were identified and investigated through colocalization analysis of GWAS, expression QTL (eQTL), and splicing QTL (sQTL), and expression analysis in blood RNA sequencing (RNA-seq). The heritability estimate for AFS was moderate at 0.276 {+/-}0.025. Three QTL regions associated with AFS were identified: Region 1 on BTA6 (43688837-45007127 bp, 1.3 Mb with three 200kb windows), explaining 0.47% of genetic variance; Region 2 on BTA6 (62770788-62967316 bp) explaining 0.27% of the genetic variance; and Region 3 on BTA18 (6299721-6498934 bp) explaining 0.17% of the genetic variance. Combining GWAS and colocalization analysis, we identified ANAPC4, RBPJ, SEPSECS, SLC34A2, ZCCHC4, CCDC149, GNPDA2, GUF1, DHX15, SOD3, GUF1 and GNPDA2 as candidate genes. These genes are enriched in signaling pathways such as progesterone-mediated oocyte maturation, parathyroid hormone synthesis, secretion and action, and selenocompound metabolism. ANAPC4, DHX15, GUF1, SEPSECS, and SOD3 were also differently expressed genes (DEGs) in blood RNA-seq. Integrating the GWAS, colocalization, differential expression analysis, and existing literature, ANAPC4 and rs136363104 were highlighted as the putative causal gene and variant for puberty, respectively. In conclusion, our study identified three candidate regions and 12 candidate genes, confirming ANAPC4 as a putative causal gene influencing AFS. This research enhances the understanding of the genetic basis of AFS in Chinese Holstein heifers. The identified key genomic regions, candidate genes, and variants have the potential to improve reproductive efficiency and reduce economic costs in heifer rearing. Interpretative summaryThe economic burden of heifer-raising and reproductive inefficiency in dairy cattle is evident. The most important strategy for optimizing heifer-raising is selecting heifers with earlier puberty (sexual maturity) to shorten age at first service (AFS). In this study, we have identified key genomic regions and candidate genes on AFS in Chinese Holstein heifers by genome-wide association study, and validated them by colocalization and RNA-seq analysis. The findings have the potential to enhance the understanding of the genetic basis of reproductive traits, leading to more targeted breeding strategies for improved reproductive efficiency.

8
Using Deep Learning with Different Architectures to Recognize RNA:DNA Triplex Structures from Histone Modification Features

Tsenum, J. L.

2025-11-24 bioinformatics 10.1101/2025.09.16.676231 medRxiv
Top 0.1%
26.9%
Show abstract

Long non-coding RNAs (lncRNAs) can perform their regulatory roles by forming triple helices through RNA-DNA interaction. Although this has been verified by few in vivo and in vitro methods, in silico approaches that seek to predict the potentials of lncRNAs and DNA sites becoming a triplex forming structure is required. Triplexator have also predicted vast amounts of lncRNAs and DNA sites that has the potentials of becoming a triplex structure. There is also an emerging experimental-evidence that the presence of epigenetic marks at DNA sites and lncRNAs can facilitate the formation of RNA:DNA triplex structures. There is therefore, a huge demand for computati onal approaches such as deep learning that can make novel predictions about RNA:DNA triplex structure formation. In this study, we developed four (4) deep neural network models that can predict the potentials of lncRNAs and DNA sites to form triple helices genome-wide, by taking histone modification marks as features. Our data was first passed through the Triplexator to screen out lncRNAs and DNA sites with low potentials of forming triple helices. We used different deep learning architectures to build our models, including two-layer convolutional neural networks (CNN) and multilayer perceptron (MLP). Our DNA2_CNN model performed best at a mean AUC of 0.78 at 32 Kernel size and learning rate of 1e-3. Our deep neural network models revealed several novel lncRNAs and DNA sites, including HOTAIR, MEG3, PARTICLE, DACOR1, MIR100HG, FENDRR, ANRIL, TUG1, MALAT1, LINC00599, TINCR, NEAT1, roX2, DHFR, OTX2-AS1, Xist, SNHG16, ATXN8OS, BCYRN1, TERC, Khps1, that have the potential of forming triplex structures, thereby confirming previous experimental results and that of the Triplexator. The performance of our models also supports previous findings that histone modification marks can help in identifying lncRNAs and DNA regions that have the potentials of forming RNA:DNA triplex structures. In conclusion, we showed that different deep learning architectures can recognize lncRNAs and DNA that have the potentials of forming RNA:DNA triplex structures.

9
Non-parametric GWAS: Another View on Genome-wide Association Study

Hu, X.; Yu, S.; Jiang, H.

2022-11-14 genetics 10.1101/2022.11.11.516099 medRxiv
Top 0.1%
26.6%
Show abstract

Genome-wide association study (GWAS) is a fundamental step for understanding the genetic link to traits (phenotypes) of interest, such as disease, BMI and height. Typically, GWAS estimates the effect of SNP on the phenotype using a linear model by coding SNP as working code, {0, 1, 2}, according to the minor allele frequency. Looking inside the linear model, we find that the coding strategy of SNP plays a key role in detecting SNPs contributed to the phenotype. Specifically, a partial mismatch between the order of the working code and that of the underlying true code will lead to false negatives, which has been ignored for a long time. Motivated by this phenomenon, we propose an indicator of possible false negatives and several non-parametric GWAS methods independent of coding strategy. Results from both simulations and real data analysis show the advantages of new methods in identifying significant loci, indicating their important complementary role in GWAS.

10
Graph Embedding Method Based Genetical Trajectory Reveals Migration History Among East Asians

Wei, Z.; Chang, C.-W.; Luo, V.; Bian, B.; Ding, X.

2019-12-10 genetics 10.1101/870253 medRxiv
Top 0.1%
23.1%
Show abstract

An important issue in human population genetics is the ancestry. By extracting the ancestral information retained in the single nucleotide polymorphism (SNP) of genomic DNA, the history of migration and reproduction of the population can be reconstructed. Since the SNP data of population are multidimensional, their dimensionality reduction can demonstrate their potential internal connections. In this study, the graph and structure learning based Graph Embedding method commonly used in single cell mRNA sequencing was applied to human population genetics research to decrease the data dimension. As a result, the human population trajectory of East Asia based on 1000 Genomes Project was reconstructed to discover the inseparable relationship between the Chinese population and other East Asian populations. These results are visualized from various ancestry calculators such as E11 and K12B. Finally, the unique SNPs along the psudotime of trajectory were found by differential analysis. Bioprocess enrichment analysis was also used to reveal that the genes of these SNPs may be related to neurological diseases. These results will lay the data foundation for precision medicine.

11
Genome-wide association study and sequence similarity analysis for unilateral renal agenesis using heterogeneous stock rats undercovers the KIT gene and AHR, ATF3, GATA3, HNF1B, POU2F2, and TFCP2 transcription factors as potential candidates to explain incomplete penetrance

Leal-Gutierrez, J. D.; Munro, D.; Chen, D.; Cheng, R.; Wang, T.; Chen, H.; Meyer, P.; Ishiwari, K.; Robinson, T.; Rau, C.; Garrett, M.

2023-09-19 genetics 10.1101/2023.09.19.558496 medRxiv
Top 0.1%
22.4%
Show abstract

1Human unilateral renal agenesis is a congenital urinary tract malformation. Affected individuals have only one kidney, which is often an asymptomatic developmental defect. A total of 5,585 male and female HS rats were assessed for unilateral renal agenesis and genotyped for 3513,321 markers. The R package SAIGEgds was used for the association analysis. The adjusted p-value threshold for the association analysis determined by permutation was equal to 5.6 (-log10). Two additional datasets were used as validation tests. Population two included 1,577 rats genotyped for 7,425,889 markers and a case-control imbalance equal to 1:174; population three included 1,407 rats, genotyped for 254,932 markers and case-control ratio equal to 1:38. The python package GxTheta was used to perform a polygenic epistasis analysis for the analyzed HS rat population. A founder haplotype mosaic determination was performed using the R package QTL2. Associated regions were selected for further analysis, including long-read PacBio sequencing for founder individuals and a founder haplotype prediction test. A similarity analysis at a genomic level and for loci encoding transcription factors predicted to interact with selected sequences inside the associated loci were accomplished. A total of 1,181 polymorphisms were associated with URA. All associated polymorphisms were located on chromosome 14 between 32.9 and 36.6 Mb. The most significant polymorphism was chr14:36,411,266, a G/T transversion. The same associated region was identified in population three. Polygenic epistasis was determined as not predominant for the presentation of URA. Based on the haplotype mosaic probability estimation, cases display a higher probability of inheriting the ACI allele. The long-read sequencing analysis showed the presence of an Erv insertion inside the intron one of the KIT gene located inside the associated region. The Erv insertion comprises one Erv sequence and two Ltr sequences located downstream and upstream of the former. No Erv insertion was identified for the founder strain BN. For ACI and HSRA, only one Ltr sequence was identified. One hundred and seven genes encoding TFs that recognize binding sites on the Erv insertion were analyzed for sequence similarity against the reference HSRA. The TF similarity score analysis for the interaction genotype and phenotype showed significance after FDR correction for 20 TFs, including AHR, HNF1B, JUNB, RARG, and RXRA. A mechanism identifying URA as a threshold phenotype is suggested in HS rats. It implies the existence of a minimum threshold for the final number of nephrons and kidney associated structures required for stalling the apoptotic process of the metanephric rudiments. Animals exhibiting a quantitative cumulative defect would express URA, being this malformation identified as a phenotype with decreased penetrance in the assessed population of HS rats. All these processes are described as mediated by KIT and TFs able to interact with sequences of the Erv insertion.

12
Abundant Parent-of-origin Effect eQTL in Humans: The Framingham Heart Study

Guan, Y.; Huan, T.; Levy, D.

2025-06-04 genetics 10.1101/2024.06.05.597677 medRxiv
Top 0.1%
22.1%
Show abstract

Parent-of-origin effect (POE) is a phenomenon whereby an alleles effect on a phenotype depends both on its allelic identity and parent from whom the allele is inherited, as exemplified by the polar overdominance in the ovine callypyge locus and the human obesity DLK1 locus. Systematic studies of POE of expression quantitative trait loci (eQTL) are lacking. In this study we use trios among participants in the Framingham Heart Study to examine to what extend POE exists for gene expression of whole blood using whole genome sequencing and RNA sequencing. For each gene and the SNPs in cis, we performed eQTL analysis using genotype, paternal, maternal, and joint models, where the genotype model enforces the identical effect sizes on paternal and maternal alleles, and the joint model allows them to have different effect sizes. We compared models using Bayes factors to identify paternal, maternal, and opposing eQTL, where paternal and maternal effects have opposite directions. The resultant variants are collectively called POE eQTL. The highlights of our study include: 1) There are more than 2, 000 genes harbor POE eQTL and majority POE eQTL are not in the vicinity of known imprinted genes; 2) Among 180 genes harboring opposing eQTL, 99 harbor exclusively opposing eQTL, and 58 of the 99 are phosphoprotein coding genes, reflecting significant enrichment; 3) Paternal eQTL are enriched with GWAS hits, and genes harboring paternal eQTL are enriched with drug targets. Our study demonstrates the abundance of POE in gene expression, illustrates the complexity of gene expression regulation, and provides a resource that is complementary to existing resources such as GTEx. We revisited two previous POE findings in light of our POE results. A SNP residing in KCNQ1 that is maternally associated with diabetes is a maternal eQTL of CDKN1C, not KCNQ1. A SNP residing in DLK1 that showed paternal polar overdominance for human obesity is a maternal eQTL of MEG3, offering an explanation for the baseline risk of homozygous samples through association between MEG3 expression and obesity. Finally, we advised caution on conducting Mendelian randomization using gene expression as the exposure.

13
blupADC: An R package and shiny toolkit for comprehensive genetic data analysis in animal and plant breeding

Mei, Q.; Fu, C.; Li, J.; Zhao, S.; Xiang, T.

2021-09-10 genetics 10.1101/2021.09.09.459557 medRxiv
Top 0.1%
21.9%
Show abstract

SummaryGenetic analysis is a systematic and complex procedure in animal and plant breeding. With fast development of high-throughput genotyping techniques and algorithms, animal and plant breeding has entered into a genomic era. However, there is a lack of software, which can be used to process comprehensive genetic analyses, in the routine animal and plant breeding program. To make the whole genetic analysis in animal and plant breeding straightforward, we developed a powerful, robust and fast R package that includes genomic data format conversion, genomic data quality control and genotype imputation, breed composition analysis, pedigree tracing, analysis and visualization, pedigree-based and genomic-based relationship matrix construction, and genomic evaluation. In addition, to simplify the application of this package, we also developed a shiny toolkit for users. Availability and implementationblupADC is developed primarily in R with core functions written in C++. The development version is maintained at https://github.com/TXiang-lab/blupADC. Supplementary informationSupplementary data are available online

14
Dopaminergic Gene Dosage in Autism versus Developmental Delay: From Complex Networks to Machine Learning approaches

Santos, A.; Caramelo, F.; Barbosa de Melo, J.; Castelo-Branco, M.

2020-04-29 genetics 10.1101/2020.04.28.065987 medRxiv
Top 0.1%
20.4%
Show abstract

The neural basis of behavioural changes in Autism Spectrum Disorders (ASD) remains a controversial issue. One factor contributing to this challenge is the phenotypic heterogeneity observed in ASD, which suggests that several different system disruptions may contribute to diverse patterns of impairment between and within study samples. Here, we took a retrospective approach, using SFARI data to study ASD by focusing on participants with genetic imbalances targeting the dopaminergic system. Using complex network analysis, we investigated the relations between participants, Gene Ontology (GO) and gene dosage related to dopaminergic neurotransmission from a polygenic point of view. We converted network analysis into a machine learning binary classification problem to differentiate ASD diagnosed participants from DD (developmental delay) diagnosed participants. Using 1846 participants to train a Random Forest algorithm, our best classifier achieved on average a diagnosis predicting accuracy of 85.18% (sd 1.11%) on a test sample of 790 participants using gene dosage features. In addition, we observed that if the classifier uses GO features it was also able to infer a correct response based on the previous examples because it is tied to a set of biological process, molecular functions and cellular components relevant to the problem. This yields a less variable and more compact set of features when comparing with gene dosage classifiers. Other facets of knowledge-based systems approaches addressing ASD through network analysis and machine learning, providing an interesting avenue of research for the future, are presented through the study. Lay SummaryThere are important issues in the differential diagnosis of Autism Spectrum Disorders. Gene dosage effects may be important in this context. In this work, we studied genetic alterations related to dopamine processes that could impact brain development and function of 2636 participants. On average, from a genetic sample we were able to correctly separate autism from developmental delay with an accuracy of 85%.

15
Plasma-free samples for transcriptomic analysis: a potential alternative to whole blood samples

Chen, Q.; Guo, X.; Wang, H.; Sun, S.; Jiang, H.; Zhang, P.; Shang, E.; Zhang, R.; Cao, Z.; Niu, Q.; Zhang, C.; Liu, Y.; Zheng, Y.; Yu, Y.; Hou, W.; Shi, L.

2023-04-28 bioinformatics 10.1101/2023.04.27.538178 medRxiv
Top 0.1%
19.6%
Show abstract

RNA sequencing (RNAseq) technology has become increasingly important in precision medicine and clinical diagnostics and emerged as a powerful tool for identifying protein-coding genes, performing differential gene analysis, and inferring immune cell composition. Human peripheral blood samples are widely used for RNAseq, providing valuable insights into individual biomolecular information. Blood samples can be classified as whole blood (WB), plasma, serum, and remaining sediment samples, including plasma-free blood (PFB) and serum-free blood (SFB) samples. However, the feasibility of using PFB and SFB samples for transcriptome analysis remains unclear. In this study, we aimed to assess the viability of employing PFB or SFB samples as substitute RNA sources in transcriptomic analysis and performed a comparative analysis of WB, PFB, and SFB samples for different applications. Our results revealed that PFB samples exhibit greater similarity to WB samples in terms of protein-coding gene expression patterns, differential expression gene profiling, and immunological characterizations, suggesting that PFB can be a viable alternative for transcriptomic analysis. This contributes to the optimization of blood sample utilization and the advancement of precision medicine research.

16
Targeted next-generation sequencing of Candidate Regions Identified by GWAS Revealed SNPs Associated with IBD in GSDs

Peiravan, A.; Salavati, M.; Psifidi, A.; Sharman, M.; Kent, A.; Watson, P.; Allenspach, K.; Werling, D.

2021-04-21 genetics 10.1101/2021.04.20.440584 medRxiv
Top 0.1%
19.3%
Show abstract

Canine Inflammatory bowel disease (IBD) is a chronic multifactorial disease, resulting from complex interactions between the intestinal immune system, microbiota and environmental factors in genetically predisposed dogs. Previously, we identified several single nucleotide polymorphisms (SNP) and regions on chromosomes (Chr) 7, 9, 11 and 13 associated with IBD in German shepherd dogs (GSD) using GWAS and FST association analyses. Here, building on our previous results, we performed a targeted next-generation sequencing (NGS) of a two Mb region on Chr 9 and 11 that included 14 of the newly identified candidate genes, in order to identify potential functional SNPs that could explain these association signals. Furthermore, correlations between genotype and treatment response were estimated. Results revealed several SNPs in the genes for canine EEF1A1, MDH2, IL3, IL4, IL13 and PDLIM, which, based on the known function of their corresponding proteins, further our insight into the pathogenesis of IBD in dogs. In addition, several pathways involved in innate and adaptive immunity and inflammatory responses (i.e. T helper cell differentiation, Th1 and Th2 activation pathway, communication between innate and adaptive immune cells and differential regulation of cytokine production in intestinal epithelial cells by IL-17A and IL-17F), were constructed involving the gene products in the candidate regions for IBD susceptibility. Interestingly, some of the identified SNPs were present in only one outcome group, suggesting that different genetic factors are involved in the pathogenesis of IBD in different treatment response groups. This also highlights potential genetic markers to predict the response in dogs treated for IBD.

17
Stratification of Systemic Lupus Erythematosus Patients Using Gene Expression Data to Reveal Expression of Distinct Immune Pathways

Deokar, A. A.

2020-10-19 rheumatology 10.1101/2020.08.25.20181578 medRxiv
Top 0.1%
19.2%
Show abstract

Systemic lupus erythematosus (SLE) is the tenth leading cause of death in females 15-24 years old in the US. The diversity of symptoms and immune pathways expressed in SLE patients causes difficulties in treating SLE as well as in new clinical trials. This study used unsupervised learning on gene expression data from adult SLE patients to separate patients into clusters. The dimensionality of the gene expression data was reduced by three separate methods (PCA, UMAP, and a simple linear autoencoder) and the results from each of these methods were used to separate patients into six clusters with k-means clustering. The clusters revealed three separate immune pathways in the SLE patients that caused SLE. These pathways were: (1) high interferon levels, (2) high autoantibody levels, and (3) dysregulation of the mitochondrial apoptosis pathway. Mitochondrial apoptosis has not been investigated before to our knowledge as a standalone cause of SLE, independent of autoantibody production, and mitochondrial proteins could be investigated as a therapeutic target for SLE in the future.

18
Genomic convergence of locus-based GWAS meta-analysis identifies DDX11 as a novel Systemic Lupus Erythematosus gene

Saeed, M.; Ibanez-Costa, A.; Patino Trives, A. M.; Collantes Estevez, E.; Aguirre-Zamorano, M. A.; Lopez-Pedrera, C.

2020-04-18 genomics 10.1101/2020.04.17.047332 medRxiv
Top 0.1%
19.2%
Show abstract

Genome-wide association studies (GWAS) of systemic lupus erythematosus (SLE) explain only [~]15% of genetic risk, indicating genes of modest effect remain to be discovered. Association clustering methods such as OASIS are more apt at identifying modest genetic effects. 410 genes were mapped to previously identified OASIS GWAS SLE loci and investigated for expression in SLE GEO datasets. GSE50395 dataset from Cordoba was used for validation. Blood eQTL for significant SNPs in SLE loci and STRING for functional pathways of differentially expressed genes was used. Confirmatory qPCR on monocytes of 12 SLE patients and controls was performed. We identified 55 genes that were differentially expressed in at least 2 SLE GEO datasets with all probes directionally aligned. DDX11 was downregulated in both GEO (P=3.60E-02) and Cordoba (P=8.02E-03) datasets and confirmed by qPCR (P=0.001). The most significant SNP, rs3741869 (P=3.2E-05) in OASIS locus 12p11.21, containing DDX11, was a cis-eQTL regulating DDX11 expression (P=8.62E-05). DDX11 interacted with multiple genes including STAT1/STAT4. Genomic convergence with OASIS and multiple expression datasets identifies novel genes. DDX11, RNA helicase involved in genome stability, is repressed in SLE. Summary StatementMore than 100 genes for SLE have been identified but they explain only [~]15% of heritability. GWAS are challenged by risk genes of modest effect. Using locus-based GWAS mapping and multiple gene expression replications, DDX11 was identified as a novel SLE gene. DDX11, repressed in SLE, may be used as a clinical diagnostic tool.

19
Analysis of 8839 pan-primate retroviral LTR elements with regulatory functions during human embryogenesis reveals their global impacts on evolution of Modern Humans.

Glinsky, G.

2023-08-07 genomics 10.1101/2023.08.06.552206 medRxiv
Top 0.1%
19.0%
Show abstract

During millions years of primate evolution, two distinct families of pan-primate endogenous retroviruses, namely HERVL and HERVH, infected primates germline, colonized host genomes and evolved to contribute to creation of the global retroviral genomic regulatory dominion (GRD) operating during human embryogenesis. Retroviral GRD constitutes of 8839 highly conserved LTR elements linked to 5444 down-stream target genes forged by evolution into a functionally-consonant constellation of 26 genome-wide multimodular genomic regulatory networks (GRNs) each of which is defined by significant enrichment of numerous single gene ontology-specific traits. Locations of GRNs appear scattered across chromosomes to occupy from 5.5% to 15.09% of the human genome. Each GRN harbors from 529 to 1486 human embryo retroviral LTR elements derived from LTR7, MLT2A1, and MLT2A2 sequences that are quantitatively balanced according to their genome-wide abundance. GRNs integrate activities from 199 to 805 down-stream target genes, including transcription factors, chromatin-state remodelers, signal sensing and signal transduction mediators, enzymatic and receptor binding effectors, intracellular complexes and extracellular matrix elements, and cell-cell adhesion molecules. GRNs compositions consist of several hundred to thousands smaller gene ontology enrichment analysis-defined genomic regulatory modules (GRMs), each of which combines from a dozen to hundreds LTRs and down-stream target genes. Overall, this study identifies 69,573 statistically significant retroviral LTR-linked GRMs (Binominal FDR q-value < 0.001), including 27,601 GRMs validated by the single ontology-specific directed acyclic graph (DAG) analyses across 6 gene ontology annotations databases. These observations were corroborated and extended by execution of a comprehensive series of Gene Set Enrichment Analyses (GSEA) of retroviral LTRs down-stream target genes employing more than 70 genomics and proteomics databases, including a large panel of databases developed from single-cell resolution studies of healthy and diseased humans organs and tissues. Genes assigned to distinct GRNs and GRMs appear to operate on individuals life-span timescale along specific phenotypic avenues selected from a multitude of down-stream gene ontology-defined and signaling pathways-guided frameworks to exert profound effects on patterns of transcription, protein-protein interactions, developmental phenotypes, physiological traits, and pathological conditions of Modern Humans. GO analyses of Mouse phenotype databases and GSEA of the MGI Mammalian Phenotype Level 4 2021 database revealed that down-stream regulatory targets of human embryo retroviral LTRs are enriched for genes making essential contributions to development and functions of all major tissues, organs, and organ systems, that were documented by numerous developmental defects in a single gene KO models. Genes comprising candidate down-stream regulatory targets of human embryo retroviral LTRs are engaged in protein-protein interaction (PPI) networks that have been implicated in pathogenesis of human common and rare disorders (3298 and 2071 significantly enriched records, respectively), in part, by impacting PPIs that are significantly enriched in 1783 multiprotein complexes recorded in the NURSA Human Endogenous Complexome database and 6584 records of virus-host PPIs documented in Virus-Host PPI P-HIPSTer 2020 database. GSEA-guided analytical inference of the preferred cellular targets of human embryo retroviral LTR elements supported by analyses of genes with species-specific expression mapping bias in Human-Chimpanzee hybrids identified Neuronal epithelium, Radial Glia, and Dentate Granule Cells as cell-type-specific marks within a Holy Grail sequence of embryonic and adult neurogenesis. Observations reported in this contribution support the hypothesis that evolution of human embryo retroviral LTR elements created the global GRD consisting of 26 gene ontology enrichment-defined genome-wide GRNs. Decoded herein the hierarchical super-structure of retroviral LTR-associated GRD and GRNs represents an intrinsically integrated developmental compendium of thousands GRMs congregated on specific genotype-phenotypic trait associations. Many highlighted in this contribution GRMs may represent the evolutionary selection units driven by inherent genotype-phenotype associations affecting primate species fitness and survival by exerting control over mammalian offspring survival genes implicated in reduced fertility and infertility phenotypes. Mechanistically, programmed activation during embryogenesis and ontogenesis of genomic constituents of human embryo retroviral GRD coupled with targeted epigenetic silencing may guide genome-wide heterochromatin patterning within nanodomains and topologically-associated domains during differentiation, thus affecting 3D folding dynamics of linear chromatin fibers and active transcription compartmentalization within interphase chromatin of human cells.

20
Immunophenotype signatures in acute leukemias unveiled by integrative systems immunology

Bahia, I. A. F.; Lima, R. D.; Oliveira, G. H. d. M.; Neta, A. P. R.; Filgueiras, I. S.; Marques, L. S.; Marques, A. H.; Fonseca, D. L. M.; Barcelos, P. M.; Nobile, A. L.; Adri, A. S.; Usuda, J. N.; Ochs, H. D.; Dias, H. D.; Nakaya, H. I.; Barroso, R. d. S.; Luchessi, A. D.; Marques, O. C.; Junior, G. B. C.

2024-06-22 oncology 10.1101/2024.06.19.24309033 medRxiv
Top 0.1%
19.0%
Show abstract

Acute leukemias (ALs) are complex hematological disorders, and accurate diagnosis is crucial for guiding treatment decisions and predicting patient outcomes. While changes in cell marker levels are well documented, the impact of these changes on marker relationships through an integrative systems approach remains uncharacterized. To address this gap, we conducted a 12-year study investigating 41 markers, including ontogenic markers and those used to diagnose both common and rare leukemia types, using immunophenotyping flow cytometry (IFC) data from 1,069 leukocyte samples obtained from peripheral blood (PB) or bone marrow (BM) aspirates of patients with suspected ALs. Machine learning techniques, such as principal component analysis (PCA) and random forest (RF) classification, demonstrated the stratification power of the cellular markers. Hierarchical clustering analysis of leukocyte ontogenetic markers revealed disease-specific clusters, irrespective of sex or sample type (PB or BM). Additionally, we found that patients with acute myeloid leukemia (AML) showed mild disruption in cell marker correlations, whereas the most significant dysregulation was observed in patients with T-cell acute lymphoblastic leukemia (T-ALL). Importantly, we identified ontogenic correlation changes indicating clusters of immature versus mature leukocyte markers, as well as cell lineage-specific markers influencing cellular relationships. These findings underscore the value of integrating systems strategies into conventional IFC analyses to enhance synthetic diagnosis and deepen our understanding of ALs pathophysiology.